Skip to main content

Backtest Metrics

The framework records more than 80 scalar measurements for each backtest run, covering returns, risk, drawdowns, trades, exposure, and execution quality. It also stores chart-ready series and a scalar summary for comparing performance across windows.

Metrics are calculated at two levels:

LevelObjectUse it for
Per runBacktestMetricsInspecting one backtest window in detail.
Per study and engineBacktestSummaryMetricsRanking strategies and judging consistency across windows.

Both are persisted in the Open Backtest Format. Opening an .obtf bundle restores the recorded values without recomputing them.

Access metrics​

Results belong to a named study and an engine slot. Select those explicitly when a bundle contains more than one study or both vector and event results:

results = app.run_backtest(strategy=strategy, study=study)
backtest = next(results.iter_backtests())

completed_study = backtest.get_study(study.name)
summary = completed_study.get_summary(engine="vector")

for run in completed_study.get_runs(engine="vector"):
metrics = run.backtest_metrics
print(
run.backtest_start_date,
metrics.total_net_gain_percentage,
metrics.sharpe_ratio,
metrics.max_drawdown,
)

Retrieve one known window directly through the study:

metrics = completed_study.get_metrics(
engine="vector",
backtest_window=completed_study.backtest_windows[0],
)

For compatibility with single-study workflows, Backtest also provides get_backtest_metrics(...) and get_all_backtest_metrics().

Value conventions​

  • Percentage fields are decimals: 0.12 means 12%.
  • Currency values use the universe's trading symbol, such as EUR or USD.
  • Durations on per-run metrics are hours unless noted otherwise.
  • Ratios such as Sharpe, Sortino, Calmar, Omega, and profit factor are dimensionless.
  • A missing or inapplicable summary value is None; for example, portfolio CAGR is unavailable for independently reset windows without an aggregate equity path. Per-run producers may retain metric-specific NaN or zero conventions.
  • total_growth and total_growth_percentage are legacy aliases for total_net_gain and total_net_gain_percentage.

Per-run metrics​

One BacktestMetrics object is attached to every completed run. The tables below cover all current fields and derived identity properties.

Identity and portfolio value​

FieldsMeaning
backtest_windowPersisted window definition that produced the run.
backtest_start_date, backtest_end_date, backtest_date_range_name, window_roleDerived identity of the active train or test range.
total_number_of_daysCalendar duration of the active range.
initial_unallocated, final_valueStarting cash and final portfolio value.
metadataFree-form run metadata.

Returns and risk-adjusted performance​

FieldsMeaning
total_net_gain, total_net_gain_percentageNet change in portfolio value.
total_growth, total_growth_percentageLegacy aliases for total net gain.
total_loss, total_loss_percentageGross loss magnitude and its share of starting capital.
gross_profit, gross_lossSum of winning and losing trade P&L.
cumulative_return, cagrTotal compounded return and annualized growth.
annual_volatilityAnnualized variability of periodic returns.
sharpe_ratioExcess return relative to total volatility.
sortino_ratioExcess return relative to downside volatility.
calmar_ratioCAGR relative to maximum drawdown.
omega_ratioProbability-weighted gains relative to losses.
profit_factorGross profit divided by gross loss magnitude.
ulcer_indexDepth and duration of percentage drawdowns.
var_95, cvar_95Value at Risk and average loss beyond VaR at 95% confidence.

Drawdown and time series​

FieldsMeaning
max_drawdown, max_drawdown_absoluteLargest peak-to-trough loss as a decimal and currency value.
max_daily_drawdownLargest one-day decline.
max_drawdown_durationDuration of the longest maximum-drawdown episode.
twr_max_drawdown, twr_max_drawdown_durationDrawdown after removing the effect of external cash flows.
equity_curve, cumulative_return_seriesPortfolio value and cumulative return over time.
rolling_sharpe_ratio, drawdown_seriesRolling risk-adjusted return and drawdown over time.
twr_equity_curve, twr_drawdown_seriesTime-weighted series that remove deposits and withdrawals.
monthly_returns, yearly_returnsCalendar return series used by reports and heatmaps.

Use the time-weighted fields when comparing portfolios with different deposit or withdrawal schedules. Raw equity and drawdown fields remain useful for account-value reporting.

Trade counts and direction​

FieldsMeaning
number_of_trades, number_of_trades_opened, number_of_trades_closed, number_of_trades_open_at_endOverall trade activity and end state.
number_of_positive_trades, number_of_negative_tradesWinning and losing closed trades.
percentage_positive_trades, percentage_negative_tradesWinner and loser shares.
number_of_long_trades, number_of_long_trades_closedLong trades opened and closed.
number_of_winning_long_trades, number_of_losing_long_trades, long_win_rateLong-side outcomes.
number_of_short_trades, number_of_short_trades_closedShort trades opened and closed.
number_of_winning_short_trades, number_of_losing_short_trades, short_win_rateShort-side outcomes.
best_trade, worst_tradeFull trade objects for the strongest and weakest outcomes.

Trade returns, timing, and streaks​

FieldsMeaning
average_trade_sizeMean trade notional.
average_trade_return, average_trade_return_percentageMean P&L across closed trades.
median_trade_return, median_trade_return_percentageMedian P&L across closed trades.
average_trade_gain, average_trade_gain_percentageMean winning-trade P&L.
average_trade_loss, average_trade_loss_percentageMean losing-trade P&L.
average_trade_duration, average_win_duration, average_loss_durationMean holding periods overall, for winners, and for losers.
current_average_trade_gain, current_average_trade_gain_percentageRecent average gain.
current_average_trade_loss, current_average_trade_loss_percentageRecent average loss.
current_average_trade_return, current_average_trade_return_percentageRecent average return.
current_average_trade_durationRecent average holding period.
win_rate, current_win_rateOverall and recent winner shares.
win_loss_ratio, current_win_loss_ratioOverall and recent average win/loss ratios.
max_consecutive_wins, max_consecutive_lossesLongest winning and losing streaks.

Excursion, exposure, and activity​

FieldsMeaning
average_mae, average_mae_percentageMean maximum adverse excursion per trade.
average_mfe, average_mfe_percentageMean maximum favorable excursion per trade.
max_mae, max_mfeLargest adverse and favorable excursions.
mfe_mae_ratioFavorable excursion relative to adverse excursion.
cumulative_exposure, exposure_ratioCapital deployed over time and fraction of time exposed.
trades_per_year, trades_per_month, trades_per_week, trade_per_dayAnnualized and calendar-normalized trading frequency.

Calendar performance​

FieldsMeaning
percentage_winning_months, percentage_winning_yearsShare of profitable calendar periods.
average_monthly_returnMean monthly return.
average_monthly_return_winning_months, average_monthly_return_losing_monthsMean return split by profitable and losing months.
best_month, worst_month, best_year, worst_yearReturn and date for calendar extremes.

Summary metrics​

BacktestSummaryMetrics rolls up all runs in one study and engine slot. Its default mode is independent_windows: every run is a separate experiment whose portfolio resets. These statistics do not create a continuous portfolio path.

Version 2 summaries record aggregation_semantics_version, aggregation_mode, return and drawdown definitions, evaluated/expected/missing window counts, and completeness. Unversioned summaries remain readable as legacy/unknown semantics and are labelled as legacy in default tables.

Aggregate performance​

FieldsMeaning
total_net_gainSum of experiment P&L in one compatible reporting currency; not continuous-account P&L.
capital_weighted_window_return, total_net_gain_percentageSum of paired window P&L divided by paired positive initial capital. The latter is a compatibility alias. The value is unavailable when any pair is incomplete or invalid.
median_window_return, worst_window_return, best_window_returnDescriptive statistics over valid per-window decimal returns.
average_net_gain, average_net_gain_percentageDuration-weighted means across windows.
total_growth, total_growth_percentage, average_growth, average_growth_percentageLegacy growth aliases.
total_loss, total_loss_percentage, average_loss, average_loss_percentageTotal and average gross-loss magnitude.
duration_weighted_mean_window_cagr, duration_weighted_mean_window_annual_volatilityExplicitly named descriptive means of per-window annualized values.
duration_weighted_mean_window_sharpe_ratio, duration_weighted_mean_window_sortino_ratio, duration_weighted_mean_window_calmar_ratioExplicitly named descriptive means of per-window ratios.
cagr, sharpe_ratio, sortino_ratio, calmar_ratio, annual_volatilityCompatibility aliases for the corresponding duration-weighted window means; they are not portfolio metrics.
worst_window_max_drawdown, max_drawdownLargest validated nonnegative per-window drawdown magnitude. The latter is a compatibility alias.
portfolio_cagr, portfolio_sharpe_ratio, portfolio_sortino_ratio, portfolio_calmar_ratio, portfolio_annual_volatility, portfolio_max_drawdown, portfolio_var_95, portfolio_cvar_95Reserved for metrics computed from one explicitly defined portfolio path; None for independent-window summaries.
profit_factorRatio recomputed from compatible pooled gross-profit and gross-loss totals.
max_drawdown_duration, var_95, cvar_95Legacy cross-window descriptive values.

Aggregate trades and exposure​

FieldsMeaning
number_of_trades, number_of_trades_closedTotal trade activity.
average_trade_return, average_trade_return_percentageMean closed-trade P&L.
average_trade_gain, average_trade_gain_percentageMean winner.
average_trade_loss, average_trade_loss_percentageMean loser.
average_trade_duration, average_win_duration, average_loss_durationAggregate holding periods.
win_rate, current_win_rate, win_loss_ratio, current_win_loss_ratioOverall and recent win/loss quality.
max_consecutive_wins, max_consecutive_lossesLongest streaks.
cumulative_exposure, exposure_ratioAggregate capital deployment.
trades_per_year, trades_per_month, trades_per_weekNormalized trading frequency.

Cross-window robustness​

FieldsMeaning
number_of_windows, window_count_evaluatedRuns included in the summary.
window_count_expected, window_count_missing, completeEvaluation-context coverage when the expected count is known.
mean_window_duration_daysArithmetic mean of valid positive window durations in days.
number_of_windows_with_tradesWindows with at least one closed trade.
number_of_profitable_windowsWindows with positive net gain.
return_consistency, win_rate_consistency, sharpe_consistencyVariation of each measure across windows; lower is more consistent.
consistency_scoreCombined consistency score from 0 to 1; higher is better.
return_stability, win_rate_stability, sharpe_stabilityPersistence of each measure across ordered windows.
stability_scoreCombined stability score from 0 to 1; higher is better.

Consistency asks whether windows produce similar outcomes. Stability asks whether those outcomes persist in order. Use both alongside the number of profitable windows; none should replace inspection of individual runs.

Rank and filter​

BacktestIndex promotes summary values to scalar columns, allowing large result sets to be filtered without loading every bundle:

import pandas as pd

# Keep pooled rows, then require acceptable risk and repeatability.
candidates = results.filter(lambda row: (
pd.isna(row["universe_key"])
and row["summary.duration_weighted_mean_window_sharpe_ratio"] >= 1.0
and row["summary.worst_window_max_drawdown"] <= 0.20
and row["summary.consistency_score"] >= 0.70
))

leader_ids = set(
candidates.df
.sort_values(
"summary.duration_weighted_mean_window_sharpe_ratio",
ascending=False,
)
.head(20)["algorithm_id"]
)
leaders = candidates.filter(
lambda row: row["algorithm_id"] in leader_ids
)

for backtest in leaders.iter_backtests():
print(backtest.algorithm_id)

After selecting candidates, load their full bundles and inspect per-window metrics, trades, and equity curves. See Backtest Reports for interactive comparison and Backtest Storage for SQLite indexing across large collections.

How the 80+ count is defined​

BacktestMetrics currently has 105 dataclass fields. The public 80+ claim is deliberately conservative: it excludes the window identity, metadata, chart series, rich trade/calendar objects, and legacy growth aliases, leaving more than 80 scalar per-run measurements. Summary fields are a separate cross-window view and are not added to that claim.